Skip to content

Gateway page: precomputed overview endpoint, faster windows, stable view - #135

Closed
bansalayush247 wants to merge 7 commits into
fedimint:masterfrom
bansalayush247:fix/gateway-page-ux
Closed

bansalayush247 wants to merge 7 commits into
fedimint:masterfrom
bansalayush247:fix/gateway-page-ux

Conversation

@bansalayush247

@bansalayush247 bansalayush247 commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

Summary

Makes the gateway page fast, current and stable. The backend now precomputes each federation's gateway list and uptime trend in the background and serves both from one endpoint. The frontend keeps its view across reloads and recovers from errors. Retired gateways no longer drag uptime down.

Changes

Backend

  • New GET /federations/:id/gateways/overview?window= returns the gateway list and the uptime trend computed at the same moment, plus computed_at. ?include=gateways or ?include=trend returns only that part.
  • A background task rebuilds the overviews for every federation, one query at a time, and keeps them in memory: 24h and 7D every 5 minutes, 30D every 30 minutes, 90D every hour (they barely change in 5 minutes and run the heaviest queries). Requests answer in ~1 ms instead of up to 4 s for 90D. Right after a restart, a missing entry is computed on the spot (and not stored, so unknown IDs can't grow the cache).
  • Removed GET /federations/:id/gateways and GET /federations/:id/gateways/uptime-trend (unstable API, only used by our frontend). /overview returns the same data.
  • Removed the 1h window: the page never offered it and it is shorter than the 5-minute poll interval.
  • Activity CTEs are NOT MATERIALIZED: the planner misestimated them as one row and nested-looped ~10k × 10k rows (90D: 55 s → 4 s, same results).
  • A gateway missing from the registry for 7 days counts as retired: it gets no more "not seen" samples, and later samples don't count against uptime or the trend. Before, every gateway ever seen stayed at 0% forever (E-Cash Club showed 28.6% while both active gateways were at 100%).

Frontend

  • One request per load, auto-refreshed every minute while the tab is visible. "Updated" shows when the server computed the data.
  • Window, status filter and sort live in the URL (refresh, Back and shared links keep them); the last window is remembered.
  • The live registry lookup runs in the background and never holds up the page.
  • Clear error and retry states, error boundaries, and a one-time reload when a lazy chunk is missing after a deploy.
  • Retired gateways are hidden by default behind a "Retired" filter.
  • Gateways count as online for 15 minutes after being seen (three polls), so they no longer flip to degraded between polls.
  • The theme is applied before first paint, so dark mode no longer flashes light on load.
  • Cards on phones; only the needed parts of ECharts are bundled.

Testing

  • cargo clippy, tsc, eslint and npm run build pass.
  • Ran the server against a local database with real federation data:
    • gateway polls are stored every 5 minutes;
    • /overview answers in ~1–2 ms for every window, including 90D;
    • each include variant returns only the requested parts, and an invalid value returns an error;
    • the gateway page loads each window with a single request and no console errors.

bansalayush247 and others added 6 commits October 9, 2026 21:27
The gateway page often showed stale data until a manual refresh, and lost
the selected time window on every reload.

- Refetch when Nginx answers from an expired cache entry (X-Cache-Status
  STALE/UPDATING), with a growing delay until fresh data arrives; refresh
  a minute after each load and on returning to the tab, never interrupting
  a load in flight; keep statuses and "seen" times ticking.
- Keep the window, status filter, sort and direction in the URL; remember
  the last window picked for visits without one.
- Render as soon as the recorded gateways arrive; run the live registry
  lookup in the background with a timeout and skip it for offline
  federations (it blocked the page for ~60 s).
- Tie each response and error to its window: a switch shows the previous
  window's numbers dimmed, a failed refresh keeps the last data with a
  Try again that shows progress, and a failed first load keeps the controls.
- Reload once per build when a lazy chunk is gone after a deploy, and add
  error boundaries so a crash can no longer blank the whole app.
- Remount federation pages per id so another federation's data never shows.
- Smaller UX fixes: no banners in normal states, stacked cards on phones,
  one accessible element per availability strip, aria-pressed controls,
  endpoints as copyable text, chart via echarts/core with dark-mode colours
  and time-zone-stable day labels, guarded localStorage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…egraded

- Apply the saved or system theme in index.html before the first paint;
  the page shipped with class="dark" and only corrected it after React
  started, so light-theme users saw a dark flash on every load.
- Count a gateway as online for up to 15 minutes since it was last seen.
  With 5-minute polls and answers cached for up to 6 minutes, the old
  10-minute limit marked healthy gateways degraded after one late poll.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ging uptime

- Mark the activity CTEs NOT MATERIALIZED. Materialized, the planner
  estimated one row for each and nested-looped ~10k x 10k rows: the 90D
  query for Global Bitcoin Federation took 55 s, now 4 s, same results.
- Treat a gateway missing from the registry for 7 days as retired: stop
  writing "not seen" samples for it, and leave samples taken after that
  out of per-gateway uptime and the trend. Previously every gateway ever
  seen counted as 0% forever (E-Cash Club showed 28.6% uptime while both
  active gateways were at 100%).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A gateway last seen more than 7 days ago has left the federation (the
same rule the backend now uses). Show it as "retired" instead of
"offline", hide it from the default list behind its own Retired filter,
and leave it out of the header counts, uptime and observed figures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The retired-gateway filter compared last_seen with `$2 - make_interval(...)`.
Postgres can't infer $2's type there and read it as an interval, so every
poll failed with "operator does not exist: timestamp with time zone >=
interval" and rolled back: no gateway was marked seen after the upgrade.
Cast the poll time explicitly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each request ran the gateway and trend queries from scratch, 4 s for 90D,
so the page leaned on Nginx serving expired cache entries and had to
detect them and retry.

- A background task rebuilds every federation's gateways and uptime trend
  for every window every 5 minutes, one query at a time, and keeps them in
  memory. Requests answer from that in ~1 ms. A missing entry (right after
  a restart) is computed on the spot but not stored, so unknown federation
  IDs can't grow the cache.
- New GET /federations/:id/gateways/overview?window= returns both lists
  computed at the same moment, with computed_at. ?include=gateways or
  ?include=trend returns only that part.
- Remove /gateways?window= and /gateways/uptime-trend, now unused.
- The page makes one request, shows "Updated" from computed_at, and drops
  the stale-cache detection and retries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bansalayush247 added a commit to bansalayush247/fedimint-observer that referenced this pull request Oct 9, 2026
Brings the gateway page fixes and the precomputed /gateways/overview endpoint (upstream PR fedimint#135) onto master, alongside the UTXO claims work.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On staging the overview task took ~75 s of every 5-minute cycle and
raised the VPS CPU from ~27% to ~35%, slowing the other database work
too. Most of that went into 30D and 90D, which barely change in 5 minutes.

- 24h and 7D still refresh every 5 minutes, 30D every 30 and 90D every
  hour. The first cycle after a restart builds all of them.
- Remove the 1h window: the page never offered it and it is shorter than
  the 5-minute poll interval. window=1h now returns the usual error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bansalayush247

Copy link
Copy Markdown
Member Author

Superseded by #137, which keeps only the gateway fetch and gateway details page changes and also fixes the connection leak in /config/:invite/gateways. The app-wide frontend changes from this PR (route error boundary, reload after deploy, theme before first paint) are dropped from this round.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant